The Lancet Digital Health
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match The Lancet Digital Health's content profile, based on 25 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.
Show abstract
Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.
Bogle, R. G.; Bogle, C. M.
Show abstract
Background: Public and clinical attention to postural orthostatic tachycardia syndrome (POTS) has increased, particularly since the COVID-19 pandemic. We quantified changes in United Kingdom Google search interest and examined whether searches increasingly used diagnostic and self-assessment language. Methods: We extracted monthly Google Trends relative search volume (RSV; 0-100) for the Health-category search term 'Pots syndrome' in the United Kingdom from January 2004 through July 2026. Five extraction attempts were made; two returned complete, identical monthly series and were retained. Prespecified eras were summarised and an exploratory interrupted time-series model at March 2020 used ordinary least squares with Newey-West heteroskedasticity and autocorrelation consistent standard errors (12 lags). Comparator searches included conventional orthostatic diagnoses, POTS diagnostic terms, associated conditions and YouTube searches. Results: The primary series comprised 271 complete months. Mean RSV increased from 18.6 during 2015-2019 to 64.8 during 2022-2023 (3.49-fold) and remained 50.6 during January 2024-July 2026 (2.73-fold above baseline). Search interest peaked in October 2022 (RSV 100); July 2026 RSV was 57. The interrupted time-series model estimated an immediate March 2020 level increase of 21.8 points (95% CI 2.8-40.7; p=0.024), while the slope change was not statistically supported (0.069 points/month, 95% CI 0.299 to 0.438; p=0.713). Searches for 'POTS symptoms', 'POTS test' and 'POTS heart rate' increased more steeply than the general term, although low baseline volumes made fold changes unstable. Conclusions: UK Google search interest in POTS rose before 2020, increased sharply after the pandemic began, and remained substantially above its prepandemic baseline. The results demonstrate a sustained change in public attention, not disease incidence or social-media causation. The growth of symptom- and testing-oriented searches is compatible with increased diagnostic self-investigation and warrants linkage to referral, diagnosis and social-media exposure data.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
BAI, T.-C.; YEH, S.-C.
Show abstract
Foundation models for chest X-ray interpretation make it possible to adapt specialised visual representations with relatively small trainable modules. We report a retrospective study of Low-Rank Adaptation (LoRA) of Rad-DINO Vision Transformer Base with 14x14 patches (ViT-B/14) for 14-class multi-label classification on the National Institutes of Health (NIH) ChestX-ray14 dataset. The official test labels were accessed during earlier model development and configuration comparisons; consequently, every official-test result in this manuscript is explicitly descriptive and non-confirmatory. We used a patient-disjoint 90/10 split of the official trainval pool (77,988 training and 8,536 validation images) and retained the released 25,596-image test partition. The historically selected all-linear LoRA configuration with safe augmentation and g=37 produced a descriptive test macro AUROC of 0.8462 versus the frozen baseline of 0.8295. Comparisons of target modules, patch-token grids, and a Rad-DINO-specific local query head are reported as retrospective comparisons rather than unbiased model-selection evidence. A confident-learning diagnostic flagged 17,653 of 86,524 trainval images (20.4%); this is a model-based flag rate, not a ground-truth label-error rate. A separate counterfactual relabeling sensitivity analysis, which uses the same model to identify and rescore disagreements, changed the descriptive AUROC to approximately 0.9445 after 6,509 policy-defined flips. This value is not achieved model performance and is not a radiologist-audited label-quality ceiling. We provide a validation-only threshold and artifact protocol for future locked evaluation, but a genuinely untouched holdout and new locked selection are required for a confirmatory headline. The existing Zenodo record contains the 25 publication figures only.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Sanjaya, J.; Pathak, S.; Si, Y.; Haghi, M.; Kudrot, N. T.; Placencia, G.; Alaei, K.; Pishgar, M.
Show abstract
Diabetic neuropathy is associated with substantial systemic disease burden, but short-term mortality risk among affected intensive care unit (ICU) patients remains difficult to characterize. We evaluated whether temporal information from the first 24 hours of ICU care improves post-landmark mortality prediction beyond severity scores and static clinical summaries. Patients aged > 18 years with diabetic neuropathy were identified in MIMIC-IV v3.1. A 24-hour landmark was used: only patients alive and still hospitalized at 24 hours were included, and the outcome was subsequent in-hospital death. The final cohort included 1,347 patients, including 83 deaths (6.16%). Data were divided into an 80% development set and a locked 20% test set. Feature selection, hyperparameter tuning, calibration, and threshold selection were restricted to development data. Logistic regression, random forest, and XGBoost were evaluated. Random forest had the highest development cross-validated PR-AUC and was selected for interpretation. On the locked test set, random forest achieved an AUROC of 0.851 (95% CI 0.765-0.924), PR-AUC of 0.339, and Brier score of 0.051; XGBoost and logistic regression achieved AUROCs of 0.847 and 0.806. In a post hoc strictly nested analysis, adding temporal predictors increased discrimination across all three algorithms; random-forest AUROC increased from 0.815 with severity and static predictors to 0.870 with the full temporal representation. First-day temporal information therefore showed additional prognostic value, but external validation is required before clinical use.
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Dack, E.; Dai, C.; Hoppe, H.; Krueselmann, P.; Meiler, S.; Jutidamrongphan, W.; Wang, L.; Tang, K.
Show abstract
AI-assisted diagnostic tools typically act as a "second opinion," providing radiologists with a discrete prediction or probability score that can be consulted alongside clinical context. This treats AI as an independent advisor rather than a collaborative partner, leaving its reasoning largely opaque. We explore a complementary approach grounded in human-AI collaboration through visual interpretability. Specifically, we investigate (1) radiologist performance when diagnosing chest X-rays from images alone, and (2) whether deep learning-generated heatmaps can support radiologists during this diagnostic process, rather than merely validating a final answer. We developed an interactive application that enables readers to engage directly with model-generated heatmaps as they form their diagnoses, and conducted a user study to evaluate how this influences diagnostic behaviour and accuracy. Our findings offer new insights into integrating interpretable, spatially grounded AI feedback into radiologist workflows. Code, datasets, and the application can be found at https://github.com/eedack01/heatmap_assisted_diagnosis.
Li, Z.; Yagis, E.; Riad, A.; Windrath-Carr, O.; Arribas, M.; Sodiq, T.; Goldsmith, K.; Glampson, B.; Flott, K.; Haji, G.; Khan, Z.; Baker, C.; Mayer, E. K.
Show abstract
Venous thromboembolism (VTE) is a leading cause of preventable inpatient mortality, while the real-world performance of mandated risk assessment and the potential for automating using electronic health record (EHR) data remain unclear. We analysed 577,904 admissions and 726,896 VTE assessment forms across five NHS hospitals between 2015 and 2025 to evaluate assessment completion, concordance with structured EHR data, clinical validity, and feasibility of EHR-based automation assisted by machine learning. Overall completion was high (96.7%), and timely completion improved from 47.4% in 2015 to 90.5% in 2024. Agreement between forms and EHR data was good for common risk factors, but low-prevalence variables were often under-documented in the forms. Despite these discrepancies, form-derived thrombosis risk was associated with increased VTE incidence (OR 3.31, 95% CI 2.81-3.90). Machine learning models using first-14-hour EHR data achieved discrimination comparable to clinician-recorded variables (AUROC 0.709 vs 0.704), supporting real-time EHR-integrated assessment pre-population and decision support.
BAI, T.-C.; YEH, S.-C.
Show abstract
CXR report generation may require a vision-language model (VLM) to produce both textual findings and spatial bounding boxes. Generative 4B-7B VLMs can emit non-empty outputs on normal images and empty outputs on abnormal images, motivating explicit structural routing. To evaluate whether a hard inference-time gate before a probabilistic VLM changes output-presence performance and to identify the mechanisms underlying paired STRUCT outcomes. We evaluated CXRxVLM v2, combining a frozen microsoft/rad-dino ViT-B/14 encoder with a 768[->]1 logistic probe (threshold 0.0557) and google/medgemma-4b-it with the pamessina/medgemma-4b-it-cure LoRA adapter. A seed=42 stratified cohort of 500 VinDr-CXR train-pool images (250 NORMAL, 250 ABNORMAL) was compared with Lingshu-7B A_baseline and D_fewshot configurations. Exact paired McNemar tests and stratum-level output-presence analyses were prespecified for the primary configurations; MedGemma 1.5 SigLIP was exploratory. CURE achieved STRUCT = 78.0% (390/500; Wilson 95% CI 74.2-81.4), versus 73.8% for Lingshu A_baseline and 74.2% for D_fewshot. Pairwise p-values were 0.0778, 0.1042, and 0.8642. The paired decomposition showed CURE ABNORMAL non-empty-output advantage of +13.6 percentage points versus Lingshu A (p = 0.0012; +14.0 points versus D, p = 0.0007), while Lingshu had higher NORMAL empty-output rates (+5.2 to +6.4 points; p = 0.0106 and p = 0.0004). The full pipeline used 8.87 GB VRAM and 4.92 s/image mean latency; 53% of records used a 25.7 ms warm gate-negative path after model loading. Equivalent overall STRUCT scores concealed two mechanistically different output regimes: CURE favored ABNORMAL non-empty outputs, whereas Lingshu favored NORMAL empty outputs. This paired decomposition, rather than the aggregate score alone, characterizes how hard-gated and probabilistic systems route output presence.
Yi, J.; Patel, K. K.; Miller, R. J. H.; Marcinkiewicz, A. M.; Kamagate, A.; Shanbhag, A.; Hijazi, W.; Lemley, M.; Zhou, J.; Liang, J. X.; Ramirez, G.; Mostafavi, S.; Urs, M.; Spielvogel, C. P.; Slipczuk, L.; Travin, M.; Alexanderson, E.; Caraval-Juarez, I.; Packard, R. R.; Al-Mallah, M.; Ruddy, T. D.; Einstein, A. J.; Feher, A.; Miller, E. J.; Acampa, W.; Knight, S.; Le, V. T.; Mason, S.; Calsavara, V. F.; Chareonthaitawee, P.; Wopperer, S.; Kwan, A. C.; Wang, L.; Li, D.; Fishman, E. K.; Lopez-Ramirez, F.; Berman, D. S.; Kwiecinski, J.; Dey, D.; Di Carli, M. F.; Slomka, P.
Show abstract
Background: Body composition is recognized as a major determinant of health outcomes, but its multidimensional nature makes clinical adoption challenging. We sought to develop and validate a body composition index (BCI) for all-cause mortality risk assessment, integrating variables of six body composition tissues. Methods: We analyzed 28509 consecutive patients undergoing myocardial perfusion imaging with routine low-dose chest CT attenuation correction (CTAC) scans acquired during myocardial perfusion imaging (MPI) at 12 centers across four countries. An artificial intelligence-based BCI was developed in a cohort of 15037 patients CTACs by integrating the CT-derived metrics of bone, skeletal muscle, and four adipose tissue compartments, coronary artery calcium score, and basic demographic variables (age, sex, BMI). The performance of BCI for mortality prediction was validated in an internal cohort of 6444 patients and an external cohort of 7028 patients by prognosis, calibration, net benefit, and explainability. Model-based simulation of tissue metrics modification was performed to evaluate estimated mortality risk reduction. Findings: During a median of 3.5 (IQR [1.9, 5.1]) years, 4697 (16%) patients died. In the external testing cohort, the BCI demonstrated excellent discrimination for mortality (area under receiver operating characteristic curve 0.78 (95% CI [0.76, 0.79]) and Harrell concordance index 0.75 [0.73, 0.76]), calibration, and net benefit overall and across pre-specified subgroups stratified by patient characteristics and imaging protocols. Visceral adipose tissue attenuation was the most influential body composition measure, followed by skeletal muscle volume. Simulated improvement in body composition was associated with significant mortality risk reduction. Interpretation: An index combining six body composition measures obtained opportunistically from routine chest CT provides robust mortality risk stratification. By converting complex body composition information into a single interpretable score, the BCI can facilitate clinical implementation of opportunistic CT biomarkers and guide individualized preventive strategies.
Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O.; Klang, E.; Barash, Y.
Show abstract
Background Large language models are increasingly proposed to post-edit decoded text in communication brain-computer interfaces and augmentative communication. A fluent model can substitute a different intent than attempted (intent drift). Whether meaning survives or confidence flags failure is unmeasured. Methods In-silico benchmark of 20 open-weight models post-editing text (4,252,326 labeled generations) corrupted with an empirical P300 confusion matrix at five levels (0-40% character error rate, CER) across the ALS message-banking vocabulary (AUTH), a message-critical probe set, and matched controls. Outputs were scored faithful, degraded, or drift by an ensemble benchmarked against physicians. A substudy re-ran 562 messages under six interface policies (seven-model panel). Findings Detected drift rose steeply with corruption in all three corpora, from 2.2% to 60.3% at 0-40% target CER in AUTH, a stress-test upper bound, not an expected clinical rate (odds ratio 2.30 per 10-percentage-point rise in target CER). Stated confidence discriminated faithful outputs reasonably well (AUROC 0.83, 0.80-0.85) but was poorly calibrated (expected calibration error 0.32, 0.27-0.37): 28.4% of outputs at confidence 90 or higher were not faithful. Message-critical content carried a small excess after matching, surviving detector removal (rule-free OR 1.10). The ratio of faithful rescues to fluent errors exceeded 1 at low corruption but fell below 1 at 20-30% target CER. No interface policy removed drift: conservative editing and abstention lowered it, alternatives and expansion raised it; the best drifted on 18.0 per 100. A 2,281-item panel (16 of 20 models) gave moderate ensemble-versus-consensus agreement (kappa 0.41); correction lowered pooled drift 31.4% to 28.3%, and a CER-stratified physician-corrected re-analysis confirmed the dose-response at each level. Interpretation Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed. This does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed. Funding: A.G. and E.K. were supported in part by the Clinical and Translational Science Awards (CTSA) grant UL1TR002541 from the National Center for Advancing Translational Sciences, through the Harvard Catalyst | The Harvard Clinical and Translational Science Center Pilot Award Program. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Competing interests: The authors declare that they have no competing interests.
Nguyen, T. T.; Nguyen, T. D.
Show abstract
Background and Objectives: Subjective cognitive decline (SCD), self-reported worsening confusion or memory over the past year, is a common early marker of cognitive concern with relevance for Alzheimer's disease prevention and population health. Population-based machine learning benchmarks that respect temporal drift in public health surveillance remain limited. We developed a reusable multi-language prediction and interpretability framework for SCD using Behavioral Risk Factor Surveillance System (BRFSS) Cognitive Decline data. Methods: We analyzed pooled national (n = 298,944) and New York (n = 30,366) cohorts with chronological train (2015-2019), validation (national: 2020-2022; New York: 2020-2021), and locked test (2023-2024) splits. Nested LASSO identified stable predictors. Sixteen machine learning algorithms were compared under year-grouped cross-validation with SMOTE restricted to training folds. Four end-to-end Python/R pipelines (single-model or soft-voting) used validation-only isotonic calibration and Youden thresholding. Primary reporting pipelines were prespecified before test unlock (national: R tidymodels single-model; New York: Python single-model); algorithms within each pipeline were chosen by validation ROC-AUC. Post-hoc GLMs (national unweighted; New York design-weighted) and two training-only knowledge-graph layers supported interpretability. Results: Locked-test discrimination was consistent across implementations (ROC-AUC approximately 0.76-0.77). Prespecified pipelines achieved test ROC-AUC 0.770 (95% CI 0.767-0.773) nationally (R gradient boosting) and 0.762 (95% CI 0.746-0.777) in New York (Python AdaBoost). Soft-voting pipelines performed similarly (national 0.770; New York 0.757) and were treated as sensitivity benchmarks. Predicted probabilities were reasonably calibrated (Brier 0.118 nationally; 0.112 in New York), and higher scores among SCD-positive respondents persisted across survey years. Difficulty deciding, mental health, and functional health items ranked highest across permutation importance, SHAP, and GLMs. Respondents who reported no difficulty deciding (DECIDE = 2) had substantially lower odds of SCD than those who reported difficulty (DECIDE = 1; aOR approximately 0.13; FDR < 0.05). Training-only knowledge graphs likewise placed difficulty deciding nearest to SCD in both cohorts. Conclusions: A temporally locked, multi-pipeline BRFSS benchmark yields stable future-year SCD risk ranking, usable calibrated probability scores that remain separated by SCD status across survey years, and convergent interpretability signals. The open implementation supports reproducible surveillance-oriented machine learning for cognitive health.
Alwakeel, M.; Zaveri, S.; Buck, E.; Rajagopal, S.; Verma, D.; Loriaux, D.; Henao, R.; Tapson, V. F.; Ortel, T. L.; Jones, W. S.; Martin, J. G.; Haines, K. L.; Freeman, N. L.; Wong, A.-K. I.
Show abstract
Background: The 2026 American Heart Association/American College of Cardiology (AHA/ACC) guidelines replaced the 2019 European Society of Cardiology (ESC) four-tier pulmonary embolism (PE) risk scheme with five clinical categories (A-E) and subcategories. These categories were set by expert consensus and have not been validated against outcomes. How patients are reclassified relative to ESC, or how the two systems compare prognostically, is unknown. Methods: We utilized three cohorts of patients with confirmed PE using structured electronic health record data, laboratory biomarkers, and large-language-model abstraction of radiology reports: Duke University Health System (n=12,992, drawn from 95,760 consecutive inpatient CT pulmonary angiography studies, 2014-2025, with no referral or registry enrollment step between imaging and cohort entry), INSPECT (Stanford; n=3,870), and MIMIC-IV (Beth Israel Deaconess; n=361). Patients were assigned AHA/ACC categories B through E, subcategorized where data allowed, and mapped to 2019 ESC risk strata. The primary outcome was 30-day mortality; discrimination was assessed with Harrell C-index. Results: Among 17,223 patients with confirmed PE, pooled 30-day mortality rose monotonically across categories: 1.5% (B), 8.9% (C), 15.5% (D), and 31.9% (E), with the ordering preserved in all three cohorts despite differing baseline mortality. Subcategory-level discrimination was reliable only at the high-acuity extreme (D2-E2); across subcategories C1 through D1, mortality did not order monotonically (9.2%, 10.8%, 8.1%, 10.9%), and adding subcategories to category C did not improve discrimination at Duke (C-index 0.699 vs 0.699). Category C patients lacking both echocardiography and biomarker testing (12.7% of category C) had mortality (10.4%) equal to or exceeding classified peers. Relative to ESC, the frameworks were concordant at the extremes, but 5.7%of ESC intermediate-risk patients were reclassified to category D, with modestly higher but non-significant 30-day mortality than those remaining in category C (10.8% versus 8.9%). Conclusions: Across a three-health-system cohort, the 2026 AHA/ACC framework produced a reproducible mortality gradient at the category level, with added subcategory granularity refining risk chiefly at the highest-acuity tiers. Discrimination across the broad intermediate band was limited, and reclassification from ESC fell almost entirely within this range.
LEI, P.; XU, Y.; ZHANG, Y.
Show abstract
Background: The condition of a patient with acute stroke often changes within hours of ICU admission. Prognostic work here targets fixed endpoints predicted from admission data, and trajectory phenotyping assigns one label per patient. We used longitudinal ICU data to identify interpretable dynamic clinical states, characterize transitions between them, and relate the current state to later events. Methods: Retrospective cohort study of 6368 adults with acute stroke in MIMIC IV v3.1. The first 72 h were divided into twelve 6-hour windows, and a hidden Markov model was fitted to 21 neurological, physiological and organ support variables. State number was chosen against criteria fixed before fitting: statistical fit, restart stability, state occupancy and clinical interpretability. Generalized estimating equations related the current state to new mechanical ventilation and vasopressor use within 12 h, and to ICU death within 72 h. Eleven sensitivity analyses assessed the robustness of the state solution. Results: Four states were selected: neurologically preserved-low support, neurological impairment low support, impairment renal dysfunction and impairment-respiratory support (63.3%, 7.8%, 11.8% and 17.1% of windows). Within 72 h, 40.3% of patients changed state at least once, and transitions ran in both directions rather than along a single severity gradient. States were identified without outcome data, yet ICU mortality by last state ranged from 2.9% to 43.9%. Adjusted for age, sex, subtype and Charlson index, the current state remained associated with organ-support escalation and death. State prevalence differed by at most 1.1 percentage points between training and test sets, and 10 of 11 sensitivity analyses gave a stable four-state solution (ARI 0.754 0.955). Conclusions: The early ICU course of acute stroke can be represented as movement among a small number of clinically interpretable states. The representation was reproducible in a held out set and across admission eras, but requires validation in an independent database before any clinical use.
Kim, S. H.; Le Guellec, B.; Rossmueller, P.; Schramm, S.; Boese, L.; Nikoubashman, O.; Kottlors, J.; Lichtenstein, T.; Strotzer, Q.; Meddeb, A.; Ziegelmeyer, S.; Steinhelfer, L.; Prucker, P.; Berberich, C.; Canisius, J.; Kreutzinger, V.; Hartl, F.; Schmitzer, L.; Rosenkranz, E.; Leonhardt, Y.; Beutel, T.-M.; Bitzer, F.; Maegerlein, C.; Boeckh-Behrens, T.; Baum, T.; Makowski, M. R.; Kirschke, J. S.; Bressem, K. K.; Adams, L. C.; Baird, G. L.; Wiestler, B.; Hedderich, D. M.
Show abstract
Background Even a highly accurate diagnostic test can yield more false-positive than true-positive findings in low-prevalence settings, which is known as the false positive paradox. Radiologists' unawareness of this paradox may foster automation bias, the tendency to excessively rely on AI outputs. Methods In this prospective, multinational, randomized controlled reader study (DRKS00038740), 34 readers from 10 countries (16 residents, 8 general radiologists or fellows, and 10 neuroradiologists) were randomly assigned to a control group (n = 17) or intervention group (n = 17), stratified by experience level. The intervention group reviewed a short, 3-minute educational video explaining the false positive paradox prior to the reading session. Both groups evaluated 20 TOF-MRA studies with AI-flagged findings (10% true-positive, 90% false-positive). Primary outcomes were acceptance rate of false-positive AI findings and follow-up intensity. These were evaluated using mixed models with crossed random effects for reader and case. Results At baseline, readers vastly overestimated the positive predictive value of AI tools for intracranial aneurysm detection (mean estimate, 62.9%; simulation-based estimate, 15.4% [95% interval, 8.1-28.0%]). The intervention reduced the odds of accepting AI false positives (OR 0.50 [upper 95% confidence bound, 0.95], one-sided p = 0.017), with acceptance probabilities of 12.7% (95% CI, 6.0-25.0%) in the intervention group compared to 22.5% (95% CI, 11.6-39.2%) in the control group. The intervention group exhibited a downward shift in follow-up intensity for false positives (OR 0.47 [upper 95% confidence bound, 0.81]; one-sided p = 0.014), recommending follow-up in 39.2% (120/306) of cases, compared to 54.9% (168/306) in the control group. Conclusion A brief education on the false positive paradox improved trust calibration in AI-assisted intracranial aneurysm detection. Our findings highlight the potential of reader-side cognitive debiasing strategies to improve trust calibration and support safer use of AI in radiology.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Nielsen, M.; Castelo, A.; Altaie, M.; Bennett, J.; Anthony, A.; Siddiqi, N. S.; Gupta, A. C.; Brock, K. K.; Woodland, M.
Show abstract
Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.
Mostafavi, S.; Shanbhag, A.; Ramirez, G.; Lemley, M.; Miller, R. J. H.; Chareonthaitawee, P.; Liang, J. X.; Dey, D.; Kavanagh, P. B.; Slipczuk, L.; Travin, M. I.; Alexanderson, E.; Carvajal Juarez, I.; Packard, R. R.; Al-Mallah, M. H.; Einstein, A. J.; Ruddy, T. D.; deKemp, R. A.; Boczar, K.; Feher, A.; Buechel, R. R.; Acampa, W.; Knight, S.; Le, V. T.; Rosamond, T. L.; Berman, D. S.; Di Carli, M. F.; Slomka, P.
Show abstract
Background: Positron emission tomography (PET) myocardial perfusion imaging (MPI) provides complementary information on perfusion, myocardial blood flow and ventricular function. While these markers are often considered collectively during interpretation, their quantitative integration with imaging and clinical data into a unified predictive framework remains limited. We developed a multimodal artificial intelligence framework that combines PET polar maps with quantitative imaging and clinical features to improve obstructive coronary artery disease (CAD) detection. Methods: We retrospectively analyzed the multicenter REFINE PET registry. Among 38,682 PET MPI studies from 14 sites, 2,833 patients without known prior CAD underwent invasive coronary angiography within 180 days. Obstructive CAD was defined as >=50% left main stenosis or >=70% stenosis in other major epicardial coronary arteries. We developed a two-stage contrastive learning framework to learn multimodal PET representations from studies without angiographic labels and transfer them to supervised CAD prediction. In Stage 1, PET image and tabular encoders were pretrained on 12,225 PET MPI studies from eight development sites using 15-channel PET polar maps, quantitative PET perfusion, flow and gated functional measures, and clinical variables. In Stage 2, the pretrained encoders and a lightweight classification head were fine-tuned in 968 angiography-labeled patients, using lower encoder learning rates to limit overfitting. The model was externally validated for angiographically defined obstructive CAD detection in 1,865 patients from six independent sites and compared with standard PET MPI metrics. Results: The prevalence of obstructive CAD was 60% in the training cohort (66% male, median age of 70 years [63, 77]), and 55% in the external validation cohort (64% male, median age of 67 years [60-74]). In external validation, the AI model achieved an AUC of 0.85 (95% confidence interval (CI), 0.83-0.87) for obstructive CAD detection and outperformed conventional quantitative PET metrics (all P < 0.001). At a specificity matched to visual summed stress score, the AI model achieved higher sensitivity (89% [95% CI, 87-91] versus 85% [95% CI, 82-87]) and negative predictive value (81% [95% CI, 77-84] versus 73% [95% CI, 69-77]; both p<0.001). The overall net reclassification improvement was 8.9% (95% CI, 4.2-13.6%; p = 0.001). Conclusions: Multimodal contrastive pretraining improved obstructive CAD detection from PET imaging beyond conventional perfusion-based scoring in independent multisite external validation.